Phase 4: Neural Networks Lesson 2 of 5

Training Deep Networks:
Backpropagation

You know what a neural network looks like. Now you need to understand how it actually learns. The answer is backpropagation: an algorithm that computes how to adjust every weight in the network to reduce the prediction error.

You will learn
What backpropagation computes and why it works
The vanishing gradient problem explained
Loss functions and how to pick the right one
Optimisers: SGD, Momentum, Adam
Regularisation for neural networks: Dropout and Batch Norm

The problem: how do you train 10 million parameters?

In Lesson 4.1 you built a network with three layers. That small network already had thousands of individual weight parameters. Real-world networks today have millions or billions. You cannot adjust each weight by hand or by trial and error. You need an algorithm that can figure out exactly how to nudge every single parameter to make the network better.

That algorithm is backpropagation, short for "backward propagation of errors." It was popularised in a landmark 1986 paper by David Rumelhart, Geoffrey Hinton, and Ronald Williams published in the journal Nature. While the mathematical ideas had been developed independently by others in the preceding years, the 1986 paper made the algorithm accessible and demonstrated that it could teach hidden layers to form useful internal representations, which settled a long-standing debate about whether multi-layer networks could be trained at all.

Analogy

A chef makes a dish and it tastes wrong. To fix it, the chef needs to know which ingredient was the problem and by how much. Backpropagation does exactly this for neural networks. After a bad prediction, it works backwards through every layer and calculates how much each weight contributed to the error. Then it adjusts each weight in the direction that reduces that contribution. It assigns blame precisely and efficiently.

How backpropagation works

Backpropagation has two phases. First, the forward pass: data flows forward through the network and a prediction is made. Second, the backward pass: the error is computed, and then it is propagated backwards through the network layer by layer, updating weights along the way.

Forward pass then backward pass: the two phases of training
Loss L(ŷ,y) Forward pass (prediction) Backward pass (error signal) INPUT OUTPUT

The forward pass produces a prediction (gold arrows). The loss function compares that prediction to the true label and outputs a single error number. The backward pass (dashed red) flows that error gradient back through every layer, computing how much each weight should change. Every weight gets updated by a tiny amount in the direction that reduces the error.

The mathematics of backpropagation relies on the chain rule from calculus. For each weight, you compute: "if I increase this weight by a tiny amount, how much does the overall loss change?" This rate of change is the gradient. You then move the weight slightly in the direction that decreases the loss. You do not need to know the calculus to use neural networks, but understanding that gradients are the mechanism of learning is important for diagnosing problems.

The vanishing gradient problem

When networks become very deep, a serious problem emerges: gradients can become vanishingly small as they travel backwards through many layers. Each layer multiplies the gradient by the derivative of its activation function. For sigmoid and tanh activations, this derivative is at most 0.25. Multiply a small number by 0.25 twenty times and it becomes essentially zero.

When gradients vanish, the weights in the early layers of the network receive almost no signal. They barely update. The network struggles to learn anything in its first layers. This was the key reason that very deep networks were considered impractical for much of the 1990s and 2000s.

Three developments largely solved the vanishing gradient problem:

1. ReLU activation: For positive inputs, the derivative of ReLU is exactly 1. No shrinkage. Signals can propagate back through many layers without fading, making deep networks trainable.

2. Residual connections (ResNets, 2015): Kaiming He and colleagues at Microsoft Research introduced "skip connections" that let the gradient bypass entire blocks of layers. A residual connection adds the input of a block directly to its output: output = F(x) + x. This gives the gradient a shortcut path that avoids the problem entirely. ResNets enabled networks 100 or more layers deep to train successfully.

3. Better weight initialisation: Initialising weights with the wrong scale can cause gradients to explode or vanish even before training begins. Methods like He initialisation (for ReLU networks) and Glorot/Xavier initialisation (for tanh networks) set the starting scale of weights so that signals are preserved as they pass through the network.

Loss functions: measuring what the network gets wrong

Before backpropagation can compute any gradients, it needs a number to differentiate: the loss. The loss function converts the difference between prediction and truth into a single number that can be minimised. Choosing the right loss function for your problem is not optional.

Problem type Loss function Why
Binary classification Binary cross-entropy Measures the gap between a predicted probability and a 0/1 label. The natural choice when your output is a sigmoid probability.
Multi-class classification Categorical cross-entropy Generalises binary cross-entropy to multiple classes. Used with a softmax output layer. Each class gets a probability and the loss rewards high probability on the correct class.
Regression Mean Squared Error (MSE) Penalises large errors more heavily than small ones. The standard choice for continuous outputs. Use Mean Absolute Error if large outliers should be treated more gently.

Optimisers: how to descend the gradient

Knowing the gradient tells you which direction to move each weight. The optimiser decides how far to move and how. Vanilla gradient descent (updating all weights once per pass through the entire dataset) is rarely used in deep learning because it is too slow and can get stuck.

Optimiser Key idea When to use
SGD Update weights after each mini-batch. Faster than full-batch gradient descent but noisier. Adding momentum smooths the updates and helps escape local minima. When you want fine-grained control and are willing to tune the learning rate manually. Often best for CNNs with a careful schedule.
Adam Adapts the learning rate for each parameter individually, using estimates of the first and second moments of the gradient. Introduced by Kingma and Ba in 2014. The default starting choice for most networks. Works well with little tuning. Slightly higher memory use than SGD.
AdamW Adam with decoupled weight decay regularisation. Fixes a subtle issue in how Adam applies L2 regularisation. The default for training large language models and transformers. Generally preferred over Adam when regularisation matters.
Learning rate: the most important hyperparameter

The learning rate controls how large each weight update step is. Too large and the updates overshoot the minimum and the loss oscillates or diverges. Too small and training takes an impractically long time. A common practical approach is to start with a moderate learning rate and then reduce it (using a learning rate schedule) as training progresses. The Keras default learning rate for Adam is 0.001, which is a reasonable starting point for most problems.

Reading the training curve

When you train a neural network, Keras tracks the loss on both the training set and the validation set at each epoch. Plotting these two curves over time is the single most useful diagnostic tool in deep learning. The shape of the curves tells you exactly what is happening.

Training vs validation loss: three patterns
Underfitting Good fit Overfitting Both curves high. Model too simple or training too short. Both curves falling together and levelling at similar low values. Train still falling. Val loss rises. Model memorising training data. Training loss Validation loss

Plot both curves after every training run. If they diverge (training keeps improving while validation worsens), you are overfitting. Stop training earlier or add regularisation. If both stay high, your model lacks capacity or needs more training time.

Regularisation: preventing overfitting in neural networks

Deep networks have enormous capacity. Left unconstrained on a small dataset, they will memorise training examples rather than learning generalisable patterns. Two regularisation techniques are standard practice in neural networks.

Dropout
Proposed by Srivastava et al., 2014
During each training step, randomly set a fraction of neurons to zero. Typical rates are 20-50%. This prevents neurons from co-adapting and forces the network to learn redundant, more robust representations. At inference time, all neurons are active but their outputs are scaled down by the dropout rate.
Batch Normalisation
Proposed by Ioffe and Szegedy, 2015
Normalises the output of each layer (zero mean, unit variance) within each mini-batch during training. Stabilises training, allows higher learning rates, and has a mild regularisation effect. Batch normalisation is now standard in almost every deep network architecture and is applied before or after the activation function.
Early Stopping
Monitor validation loss each epoch
Stop training when the validation loss stops improving (or begins to worsen). This prevents the model from continuing to fit the training data after it has already reached peak generalisation. Keras provides a built-in EarlyStopping callback for this.
L2 Weight Decay
Penalise large weights in the loss
Adds a penalty term to the loss that grows with the size of the weights. This discourages the network from placing extreme importance on any single feature. Sometimes called "weight decay" or "L2 regularisation" and controlled by a small constant (typically 0.0001 to 0.01).
Python training_with_callbacks.py
from tensorflow.keras import layers, callbacks
import tensorflow as tf

model = tf.keras.Sequential([
    layers.Dense(128, activation='relu'),
    layers.BatchNormalization(),        # normalise layer outputs
    layers.Dropout(0.3),               # randomly zero 30% of neurons
    layers.Dense(64,  activation='relu'),
    layers.BatchNormalization(),
    layers.Dropout(0.3),
    layers.Dense(1,   activation='sigmoid')
])

model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])

# Early stopping: stop if val_loss doesn't improve for 10 epochs
early_stop = callbacks.EarlyStopping(
    monitor='val_loss',
    patience=10,
    restore_best_weights=True   # revert to the best epoch when stopping
)

history = model.fit(
    X_train, y_train,
    epochs=200,
    batch_size=32,
    validation_split=0.15,
    callbacks=[early_stop],
    verbose=0
)

print(f"Stopped at epoch {len(history.history['loss'])}")

With restore_best_weights=True in the EarlyStopping callback, Keras automatically reverts the model to its best state before the validation loss started worsening. You ask it to train for 200 epochs but it might stop at epoch 47, having found the best generalisation well before the end.

"The practical utility of the various tricks used in training neural nets is often underestimated. Choosing good activations, initialisations, and optimisers matters as much as choosing the architecture."

Andrej Karpathy, Stanford CS231n lecture notes (widely read AI teaching resource)
Hands-on activity

Watch your network learn: plot the training curve

Training a network blindly and hoping accuracy improves is not good practice. In this activity you will visualise training in real time, diagnose the overfitting/underfitting state of your model, and use callbacks to train smarter.

01 Open the Lesson 4.2 Colab notebook. Train the network for 100 epochs without early stopping and save the history object.
02 Plot training loss and validation loss on the same graph. At what epoch does validation loss start to rise? That is your overfitting point.
03 Add Dropout(0.4) after each hidden layer. Retrain and re-plot. Does the gap between training and validation loss decrease?
04 Add the EarlyStopping callback with patience=10. How many epochs did the model actually train for? Is the final accuracy better or worse than the 100-epoch version?
05 Try changing the learning rate from 0.001 to 0.01 and then to 0.0001. Plot all three training curves together. How does learning rate affect the shape of the curve and the final result?
Your Notes
Studying independently? Write your thoughts or answers below. Notes save automatically to your browser.
Practice Notebook
Run this lesson's code live in Google Colab
All examples + challenge exercises · Free GPU included · No setup required
Open In Colab
Progress
Done with this lesson?
Mark it complete to track your progress.